Papers with automatic evaluation metrics
Copied to clipboard
| Challenge: | Existing automated evaluation metrics for machine translation are expensive and lack inter-rater reliability. |
| Approach: | They propose a task-oriented and human-centric evaluation framework for machine translation output based on professional post-e diting annotations. |
| Outcome: | The proposed framework improves translation quality and system performance and transparency . it is cost-effective, easy to use and faster to implement . |
Copied to clipboard
| Challenge: | This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field. |
| Approach: | This tutorial presents the evolution of automatic evaluation metrics to their current state . it aims to assess the extent of scientific progress made and identify areas/components that need improvement . |
| Outcome: | This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field. |
Copied to clipboard
| Challenge: | Existing automated evaluation metrics fail to consider factual correctness or are limited in their interpretability. |
| Approach: | They propose a radiology report evaluation metric that leverages natural language understanding of language models to identify and explain clinically significant errors. |
| Outcome: | The proposed method demonstrates higher correlation with expert error counts and higher alignment with expert preferences when compared to previous methods. |
Copied to clipboard
| Challenge: | a lack of comprehensive studies on evaluation metrics for text summarization hinders progress . a new study aims to improve evaluation metrics that correlate with human judgments . |
| Approach: | They propose to re-evaluate automatic evaluation metrics and share a toolkit for evaluation . they hope to promote a more complete evaluation protocol for text summarization . |
| Outcome: | The proposed evaluation metrics are inconsistent with existing evaluation protocols. |
Copied to clipboard
| Challenge: | a new dataset and evaluation benchmark for Few-shot Region-aware Machine Translation is presented . FRMT is a type of style-targeted translation that uses labeled training data to perform tasks. |
| Approach: | They propose a dataset and evaluation benchmark for Few-shot Region-aware Machine Translation. |
| Outcome: | The proposed model is based on two translations from English into Portuguese and Mandarin Chinese. |
Copied to clipboard
| Challenge: | Existing automatic evaluation metrics for open-domain dialogue systems correlate poorly with human evaluation. |
| Approach: | They propose to construct response selection test sets with well-chosen false candidates to evaluate response generation systems via response selection. |
| Outcome: | The proposed method correlates with human evaluation better than widely used metrics such as BLEU. |
Copied to clipboard
| Challenge: | a recent study focused on machine translation evaluation for low-resource languages . linguistic aspects that vary across languages are factors that will exacerbate the problem in low-source languages due to the reliance on extensive data resources. |
| Approach: | They propose to use multi-dimensional quality metrics and DA annotations to meta-evaluate MT evaluation metrics for low-resource languages. |
| Outcome: | The proposed evaluation metrics are based on human scores on the candidate translations of assamese, maithili, and Punjabi. |
Copied to clipboard
| Challenge: | Document simplification requires complex factors such as technical terminology, metaphors, and overall coherence. |
| Approach: | They propose a multi-agent framework for document simplification based on large language models that emulates the collaborative process of a human expert team through the roles played by multiple agents. |
| Outcome: | The proposed framework emulates the collaborative process of a human expert team through the roles played by multiple agents, addressing the intricate demands of document simplification. |
Copied to clipboard
| Challenge: | Existing LLM-based conversational systems do not take into account the student’s affective states. |
| Approach: | They propose an emotionally aware LLM-powered math tutor that models student emotions and maps them to relevant pedagogical strategies. |
| Outcome: | The proposed model improves student engagement and learning effectiveness by 23 points using win rate and 3 points at an overall level using DAMR scores. |
Copied to clipboard
| Challenge: | Standard language generation metrics have been shown to be ineffective for dialog evaluation. |
| Approach: | They propose an unsupervised evaluation metric for dialog that trains unsupervised models to measure several desirable qualities of dialog. |
| Outcome: | The proposed evaluation metric strongly correlates with human judgment on Topical-Chat and PersonaChat. |
Copied to clipboard
| Challenge: | Existing methods for summarization evaluations that approximate human judgments are lacking for accuracy and reliability. |
| Approach: | They propose methods for calculating confidence intervals and running hypothesis tests for correlations using bootstrapping and permutation. |
| Outcome: | The proposed methods show that the confidence intervals are wide, demonstrating high uncertainty in the reliability of automatic metrics. |
Copied to clipboard
| Challenge: | a systematic review of automatic evaluation metrics for Natural Language Generation (NLG) shows that task-agnostic metrics have a weak correlation with human . |
| Approach: | They propose a framework to assess the effectiveness of automatic metrics in three NLG tasks . they propose task-agnostic and human-aligned metrics to be used for evaluation . |
| Outcome: | The proposed framework provides access to the evaluation tools for three NLG tasks. |
Copied to clipboard
| Challenge: | Existing summarization datasets often have issues that seriously limit their usability. |
| Approach: | They propose a faster but more straightforward approach to developing summarization benchmark data . they use a protocol that hires highly-qualified contractors to read stories and write original summaries from scratch . |
| Outcome: | The proposed protocol is faster but more straightforward than scraping summaries from everyday text. |
Copied to clipboard
| Challenge: | Recent research has demonstrated that large language models (LLMs) can translate cultural elements in languages such as idioms and proverbs. |
| Approach: | They propose to use large language models to translate culturally rooted proverbs in conversation and between languages with similar cultural backgrounds to compare their results. |
| Outcome: | The proposed models can achieve good translation between languages with similar cultural backgrounds and outperform NMT models in proverb translation. |
Copied to clipboard
| Challenge: | a novel task aims to generate engaging questions from location-aware information . a lightweight model can be used to generate such questions . |
| Approach: | They propose a task to generate engaging questions from location-aware data . they represent location-based information with surrounding images and a GPS coordinate . |
| Outcome: | The proposed method outperforms baselines regarding human evaluation and evaluation metrics. |
Copied to clipboard
| Challenge: | Story generation is a challenging problem in artificial intelligence (AI) . previous work focused on learning statistical models of event sequences from large-scale text corpora . |
| Approach: | They propose to use adversarial training to generate reasonable story endings . their model includes a generator that defines the policy of generating a story ending . |
| Outcome: | The proposed model achieves better performance on the task of Story Cloze Test with an accuracy of 62.6% compared with state-of-the-art baseline methods. |
Copied to clipboard
| Challenge: | Existing automatic evaluation metrics are based on procedures that diverge from human evaluation. |
| Approach: | They propose to aggregate automatic evaluation metrics to bridge this gap . they propose to use edit-based metrics, -gram based metrics and sentence-level metrics to find the best ranking system. |
| Outcome: | The proposed method outperforms existing metrics on the SEEDA benchmark and improves edit-based metrics, -gram based metrics and sentence-level metrics. |
Copied to clipboard
| Challenge: | Empirical results show that our proposed model outperforms the state-of-the-art methods in terms of both automatic evaluation metrics and human judgment. |
| Approach: | They propose a model which uses large-scale commonsense and named entity based knowledge to ground dialogue on external knowledge and topic-specific knowledge associated with each utterance. |
| Outcome: | The proposed model outperforms the state-of-the-art methods on two benchmark datasets. |
Copied to clipboard
| Challenge: | Existing instruction generators have not been evaluated using human wayfinders . BLEU, ROUGE, METEOR and CIDEr are ineffective for evaluating grounded navigation instructions. |
| Approach: | They propose an instruction-trajectory compatibility model that operates without reference instructions to improve wayfinding performance. |
| Outcome: | The proposed model shows the highest correlation with human wayfinding outcomes when scoring individual instructions. |
Copied to clipboard
| Challenge: | Existing approaches to commonsense inference lack coverage and expressive diversity of commonsensense knowledge graphs. |
| Approach: | They propose a framework that contrasts sets of semantically similar and dissimilar events . they propose 'solar' framework that can be used to learn commonsense inference . |
| Outcome: | The proposed framework outperforms the state-of-the-art commonsense transformer on commonsensense inference by 1.84% on average among 8 metrics. |
Copied to clipboard
| Challenge: | Existing automatic dialog evaluation metrics are mostly reference-based . Existing models that measure self-reported user ratings are biased and variance among different users. |
| Approach: | They propose an automatic evaluation model that automatically cleans self-reported user ratings as it trains on them. |
| Outcome: | The proposed model achieves 89.2% accuracy in the dialog comparison task. |
Copied to clipboard
| Challenge: | Recent studies show that doctors can save significant amounts of time when using automatic note generation. |
| Approach: | They propose task-specific metrics for automatic note generation from medical conversation summarization and generation, including knowledge-graph embedding-based metrics, customized model-based measures with domain-specific weights, and ensemble metrics. |
| Outcome: | The proposed evaluation metrics are compared to existing models and can have different behaviors on different types of clinical notes datasets. |
Copied to clipboard
| Challenge: | Recent question generation approaches assume that the answer is known . however, such passages are what is being sought when verifying a claim. |
| Approach: | They propose a method that generates questions based on different focal points within a claim . they demonstrate that the method generates more relevant and informative questions . |
| Outcome: | The proposed method outperforms previous work on a fact-checking question generation dataset on measurable evaluation metrics. |
Copied to clipboard
| Challenge: | Vaccine interventions aim to answer concerns expressed about vaccination. |
| Approach: | They propose a dataset to evaluate how well responses are tailored to a common-ground opinion . they find that GPT-4-Turbo performs significantly better than others . |
| Outcome: | The proposed dataset outperforms fine tuned LLMs on the task of tailoring vaccine responses to common-ground opinions. |
Copied to clipboard
| Challenge: | Multi-modal Large Language Models have shown remarkable progress in visual contexts, yet their ability to convert visual figures into executable code remains underexplored. |
| Approach: | They propose to use a set of visual coding metrics to assess MLLMs' visual . pass rate, text-match ratio, and GPT-4V rating judgement to assess the quality of generated code and rendered images. |
| Outcome: | The proposed benchmark includes 132 high-quality matplotlib plots across six plot types, as well as 150 and 86 plots from Python’s and R’s plotly libraries respectively, totaling 368 plots. |
Copied to clipboard
| Challenge: | Existing clarification datasets with limited annotated examples do not address ambiguous phenomena. |
| Approach: | They propose a dataset that allows users to ask clarification questions using open-domain examples. |
| Outcome: | The proposed model achieves better performance than strong baselines and provides new challenges. |
Copied to clipboard
| Challenge: | Existing code-to-text generation models produce only high-level code summaries that do not capture implementation-level choices essential for these scenarios. |
| Approach: | They propose a code explanation generation task that uses code docstrings to refine models. |
| Outcome: | The proposed model can generate well-structured long docstrings comparable to human-written ones. |
Copied to clipboard
| Challenge: | Existing work on news-image captioning requires a joint understanding of image and text. |
| Approach: | They propose a Transformer model that integrates text and image modalities and attends to textual features from visual features in generating a caption. |
| Outcome: | The proposed model outperforms the state-of-the-art model and improves the quality of news-image captions. |
Copied to clipboard
| Challenge: | Existing methods for redacting offensive comments into non-offensive ones are inadequate to detect hateful content on social media platforms. |
| Approach: | They propose a method for transforming offensive comments into non-offensive ones using a Retrieve, Generate and Edit unsupervised style transfer pipeline. |
| Outcome: | The proposed method outperforms existing models on automatic metrics and human evaluations and consistently performs well on all automatic evaluation metrics. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have introduced paradigm-shifting approaches in natural language processing, yet their transformative in-context learning (ICL) capabilities remain underutilized, especially in customer service dialogue summarization. |
| Approach: | They propose a single-instance, multi-step framework that orchestrates information extraction, self-correction, and evaluation through sequential interactive generation chains. |
| Outcome: | The proposed framework outperforms existing models and prompts in the customer service dialogue summarization domain. |
Copied to clipboard
| Challenge: | Existing methods for topic-to-essay generation are insufficient for generating novel, diverse, and topic-consistent paragraph-level text with a set of topics. |
| Approach: | They propose to integrate commonsense from external knowledge base into the generator through dynamic memory mechanism and adversarial training to further improve topic-consistency. |
| Outcome: | The proposed task is more novel, diverse, and topic-consistent than existing methods in terms of both automatic and human evaluation. |
Copied to clipboard
| Challenge: | Existing evaluation metrics focus on turnlevel quality, which is not well suited for open-end dialogue tasks. |
| Approach: | They propose to measure the performance of a dialogue system by computing the distributionwise distance between its generated conversations and real-world conversations. |
| Outcome: | The proposed metrics correlate better with human judgments than existing metrics on dialogue systems. |
Copied to clipboard
| Challenge: | Generative AI has made rapid advances in multimodal understanding and code generation. |
| Approach: | They construct a first real-world benchmark for multimodal large language models that directly convert visual designs into code implementations by manually curating 484 diverse real-life webpages as test cases. |
| Outcome: | The proposed model can generate code implementations that directly render into the given reference webpages, given the screenshots as input. |
Copied to clipboard
| Challenge: | Existing methods for unsupervised text style transfer lack parallel data and difficulties in content preservation. |
| Approach: | They propose a neural approach to unsupervised text style transfer using non-parallel data. |
| Outcome: | The proposed approach can be trained end-to-end on two widely-used public datasets. |
Copied to clipboard
| Challenge: | Recent approaches to conversation response generation model speakers and utterances together but are too tailored to the speakers. |
| Approach: | They propose a new conversation model with a stochastic variable conditioned on the speakers and affects the context. |
| Outcome: | The proposed model outperforms existing models in generating appropriate conversation responses. |
Copied to clipboard
| Challenge: | Existing evaluation metrics for open-domain dialogue systems are limited by the diversity of possible outcomings. |
| Approach: | They propose a method to augment a reference set to improve reliability . they propose BLEU to measure similarity between a predicted response and a small set of references . |
| Outcome: | The proposed model improves the reliability of reference-based metrics with augmented reference sets. |
Copied to clipboard
| Challenge: | Recent question generation approaches use the sequence-to-sequence framework to optimize the log likelihood of ground-truth questions using teacher forcing. |
| Approach: | They propose to optimize for QG-specific objectives via reinforcement learning to improve question quality. |
| Outcome: | The proposed model improves the fluency, relevance, and answerability of generated questions. |
Copied to clipboard
| Challenge: | Empirical evaluation shows our model to outperform the single-hop question generation models on both automatic evaluation metrics such as BLEU, METEOR, and ROUGE and human evaluation metrics for quality and coverage of the generated questions. |
| Approach: | They propose a question-aware reward function to maximize the utilization of supporting facts in the context. |
| Outcome: | The proposed model outperforms single-hop neural question generation models on automatic evaluation metrics and human evaluation metrics for quality and coverage of the generated questions. |
Copied to clipboard
| Challenge: | Existing models generate fluent and coherent summaries, but inconsistencies can be found in generated summary. |
| Approach: | They propose to use symbolic knowledge distillation to improve the factual consistency of smaller pretrained models for dialogue summarization. |
| Outcome: | The proposed model outperforms baseline models in BART, PEGASUS, and Flan-T5 in factual consistency and accuracy. |
Copied to clipboard
| Challenge: | Natural language generation models are a key component of deep learning, says aaron eliott . he says it is crucial to develop and apply better metrics for NLG evaluation . |
| Approach: | a new open-source library for NLG evaluation is created to facilitate researchers to judge the effectiveness of their models. the framework provides a living collection of NLG metrics in a unified and easy-to-use environment. |
| Outcome: | a new open-source library for NLG evaluation aims to improve performance of models . the framework provides tools to apply, analyze, compare, and visualize the metrics . |
Copied to clipboard
| Challenge: | Existing sequence-to-sequence neural models may not be able to identify answer-relevant context words for question generation. |
| Approach: | They propose to model the unstructured sentence and the structured answer-relevant relation for question generation by combining to the point context and unstructure. |
| Outcome: | Experiments show that the proposed model improves on the unstructured sentence and the structured answer-relevant relation. |
Copied to clipboard
| Challenge: | Existing metrics for dialog evaluation are trained on human annotations, which is cumbersome to collect. |
| Approach: | They propose to use user sentiment and other information as proxy to measure the quality of previous dialogs. |
| Outcome: | The proposed model is comparable to models trained on human annotated data. |
Copied to clipboard
| Challenge: | Existing approaches to argument summarization rely on single-pass generation, offering limited support for factual correction or structural refinement. |
| Approach: | They propose a large language diffusion framework that iteratively improves argument summarization by sufficiency-guided remasking and regeneration. |
| Outcome: | Empirical results show that Arg-LLaDA surpasses state-of-the-art baselines in 7 out of 10 evaluation metrics. |
Copied to clipboard
| Challenge: | Existing methods to generate implausible stories using plots are unnatural and oversimplify the characteristics of implusible machine-generated stories. |
| Approach: | They propose to generate a more comprehensive set of implausible stories using plots . plots are structured representations of controllable factors used to generate stories . |
| Outcome: | The proposed model improves the quality of generated implausible stories using plots . it shows that the evaluation metrics trained on the generated data correlate better with human judgments compared to baselines. |
Copied to clipboard
| Challenge: | Existing models for generating mathematical word problems are lacking in educational assessment. |
| Approach: | They propose an end-to-end neural model to generate diverse mathematical word problems from commonsense knowledge graph and equations. |
| Outcome: | The proposed model outperforms the SOTA models in terms of evaluation metrics and topic relevance. |
Copied to clipboard
| Challenge: | Using word-level linguistic annotations in under-resourced neural machine translation is challenging for many languages. |
| Approach: | They propose to use word-level linguistic annotations to label source-language (SL) or target-language words to improve translation performance. |
| Outcome: | The proposed language annotations outperform part of speech and morphological description tags in the target language, while the morpho-syntactic description tags improve the grammaticality of the output. |
Copied to clipboard
| Challenge: | Existing methods for evaluating image transcreation have relied on human evaluation. |
| Approach: | They propose a suite of automatic evaluation metrics inspired by machine translation metrics . they identify cultural relevance, semantic equivalence and visual similarity as critical dimensions of image transcreation . |
| Outcome: | The proposed evaluation metrics agree with human ratings across 7 countries. |
Copied to clipboard
| Challenge: | Existing methods for document simplification address complex factors such as technical terminology, metaphors, and overall coherence. |
| Approach: | They propose a multi-agent framework AgentSimp for document simplification based on large language models that simulates collaboration among agents through roles played by multiple agents. |
| Outcome: | The proposed framework produces simplified documents that are more thoroughly simplified and more coherent across various articles and styles. |
Copied to clipboard
| Challenge: | Existing evaluation metrics are not designed to cope with this flexibility. |
| Approach: | They propose to group the qualities into three groups to obtain a single metric called USL-H. |
| Outcome: | The proposed metric achieves good correlations with human judgment and maintains its configurability towards different aspects and metrics. |
Copied to clipboard
| Challenge: | Existing studies on large language models lack adequate evaluations and prompting strategies for explainability. |
| Approach: | They evaluate the mental health analysis and emotional reasoning ability of large language models (LLMs) using 11 datasets across 5 tasks. |
| Outcome: | The proposed model shows strong in-context learning ability but still has a significant gap with advanced task-specific methods. |
Copied to clipboard
| Challenge: | Recent studies have focused on the non-deterministic properties of language models, but these properties remain under-explored in machine translation. |
| Approach: | They propose a method that evaluates MT systems and identifies temperature-constrained non-deterministic MT as a distinct phenomenon. |
| Outcome: | The proposed framework provides higher-quality candidates than Deterministic MT under temperature constraints. |
Copied to clipboard
| Challenge: | Text generation from semantic parses is challenging due to the complexity of the inner logic and the lack of automatic evaluation metrics for logic consistency. |
| Approach: | They propose a framework for logic consistent text generation from semantic parses that employs iterative training procedures and quality control. |
| Outcome: | The proposed framework enhances logic consistency and human evaluation on two benchmark datasets. |
Copied to clipboard
| Challenge: | Traditional metrics for automatic text evaluation are tailored to specific tasks, while LLM-based evaluation metrics are costly. |
| Approach: | They propose a metric that leverages projections of LLM representations for evaluation. |
| Outcome: | The proposed metric exhibits higher correlation with human judgments than previous methods on 14 datasets. |
Copied to clipboard
| Challenge: | Non-factoid (NF) question answering is challenging to evaluate due to diverse potential answers and no objective criterion. |
| Approach: | They propose a listwise NFQA evaluation approach that uses Large Language Models to rank candidate answers in a descending list of reference answers sorted by descending quality. |
| Outcome: | The proposed method has higher correlations with human annotations than standard methods. |
Copied to clipboard
| Challenge: | Existing studies on text-based QG focus on generating SQuAD-style questions. |
| Approach: | They propose a multi-hop question generation model that does context encoding in multiple hops with Graph Convolutional Network and encoder fusion via an Encoder Reasoning Gate. |
| Outcome: | Empirical results show that the proposed model generates fluent questions with high completeness and outperforms baselines on automatic evaluation metrics. |
Copied to clipboard
| Challenge: | Existing evaluation metrics based on n-gram similarity do not correlate well with human judgments . large datasets for document Question Answering (QA) have enabled the development of end-to-end supervised models . |
| Approach: | They propose a scoring function to capture answerability of questions . they also integrate existing similarity metrics into the scoring function . |
| Outcome: | The proposed scoring function improves human judgments on question answerability . the proposed scoring functions are made publicly available . |
Copied to clipboard
| Challenge: | Existing evaluation metrics for image captioning are primarily designed for short captions and are not suitable for long captions. |
| Approach: | They propose an automatic evaluation metric for long captions developed within a novel LLM-Hybrid-as-a-Judge framework. |
| Outcome: | The proposed metric outperforms existing metrics and achieves superhuman performance on LongCap-Arena. |
Copied to clipboard
| Challenge: | Existing studies on visual storytelling (VIST) use automated evaluation metrics for text generation. |
| Approach: | They develop a Vrank metric that repurposes human evaluation results for automatic evaluation. |
| Outcome: | The proposed model is more accurate than existing metrics and is generalizable to textual stories. |
Copied to clipboard
| Challenge: | Existing methods for dialog agent training lack a robust action space for entangled information, which can cause bias and deviate from natural human language. |
| Approach: | They propose phrase-level action reinforcement learning which allows the model to alter the sentence structure and content with the sequential action selection. |
| Outcome: | The proposed model achieves competitive results with state-of-the-art models on the MultiWOZ dataset, indicating that it is effective for solving task-oriented dialogs. |
Copied to clipboard
| Challenge: | In natural language, we often omit some words that are easily understandable from the context. |
| Approach: | They propose to use a dataset to evaluate whether translation models can resolve zero pronoun problems in Japanese to English translations. |
| Outcome: | The proposed model can resolve the zero pronoun problem in Japanese to English translations. |
Copied to clipboard
| Challenge: | Existing automatic evaluation metrics for summarization are insensitive to factual inconsistencies. |
| Approach: | They propose an automatic evaluation protocol that detects factual inconsistencies in a model-generated summary. |
| Outcome: | QAGS has higher correlations with human judgments of factual consistency than other evaluation metrics. |
Copied to clipboard
| Challenge: | Text simplification (RCTS) models often depend on parallel corpora with readability annotations on both source and target sides. |
| Approach: | They propose to use instruction-tuned large language models for zero-shot RCTS to reduce reliance on parallel corpora with readability annotations on both source and target sides. |
| Outcome: | The proposed model can generate sentences with the desired readability, but the model's limitations and characteristics of the source sentences impede it. |
Copied to clipboard
| Challenge: | Pretraining-based (PT) evaluation metrics are not effective for training grammatical error correction systems. |
| Approach: | They propose a pretraining-based GEC evaluation metric which only uses PT-based metrics to score the corrected parts of the system. |
| Outcome: | The proposed evaluation metric outperforms existing methods on a CoNLL14 evaluation task. |
Copied to clipboard
| Challenge: | Existing models for abstractive text summarization do not provide explicit interdocument relationships among source documents. |
| Approach: | They propose a model that uses sparse attention based on the conversational structure and a multi-task training objective that predicts metadata features. |
| Outcome: | The proposed model outperforms baseline models in terms of evaluation metrics but struggle to handle conflicts in source documents. |
Copied to clipboard
| Challenge: | Experimental results show that automatic summarization generates concise summaries that contain key ideas of source documents. |
| Approach: | They propose to use Element-aware test sets to annotate news-related reference summaries to focus on more fine-grained news elements objectively and comprehensively. |
| Outcome: | The proposed method outperforms state-of-the-art fine-tuned PLMs and zero-shot LLMs by +4.33/+4.77 on the two datasets, respectively. |
Copied to clipboard
| Challenge: | Existing approaches to compose Ci are limited in handling the constraints of tune patterns . authors propose a non-autoregressive approach to generate Ci using a synchronous process . |
| Approach: | They propose to compose Ci using a non-autoregressive approach that takes into account rigid formats . they propose to apply reinforcement learning to the generation process with rigid constraints . |
| Outcome: | The proposed method outperforms baselines and previous studies on a Ci dataset . it allows the model to perform synchronous generation while maintaining the format and content requirement. |
Copied to clipboard
| Challenge: | Existing evaluation metrics are compared based on their ability to correlate with humans, but they disagree in the higher-scoring range in which current systems operate. |
| Approach: | They show that evaluation metrics which behave similarly on these datasets strongly disagree in the higher-scoring range in which current systems operate. |
| Outcome: | The evaluation metrics which behave similarly on these datasets strongly disagree in the higher-scoring range in which current systems operate. |
Copied to clipboard
| Challenge: | Recent years have brought about interest in the task of summarizing conversation threads. |
| Approach: | They develop an email thread summarization dataset that contains human-annotated short and long email threads over a wide variety of topics. |
| Outcome: | The proposed dataset contains human-annotated short (30 words) and long (100 words) summaries of 2,549 email threads over a wide variety of topics. |
Copied to clipboard
| Challenge: | Existing personalized dialogue models use human designed persona descriptions to improve dialogue consistency. |
| Approach: | They propose to extend Model-Agnostic Meta-Learning (MAML) to personalized dialogue learning without using persona descriptions. |
| Outcome: | The proposed model outperforms baseline models in terms of human-evaluated fluency and consistency on a persona-chat dataset. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) evaluation is a patchy and inconsistent landscape . established automatic evaluation metrics are poor surrogates, correlating weakly with human judgement. |
| Approach: | They propose to use both automatic and human evaluation to evaluate generative LLMs on three NLP benchmarks: text summarisation, text simplification and grammatical error correction. |
| Outcome: | The proposed model outperforms many popular models according to human reviewers on the majority of metrics, while scoring much worse when using classic automatic evaluation metrics. |
Copied to clipboard
| Challenge: | Existing methods for generating MWP text from equations are inflexible and require pre-defined templates. |
| Approach: | They propose a neural model which generates MWPs from equations by constructing a Quantity Cell Graph from the retrieved MWp instance and reasoning over it. |
| Outcome: | The proposed model performs impressively on educational MWP set and on human evaluation metrics. |
Copied to clipboard
| Challenge: | Current state-of-the-art neural dialogue models learn from human conversations . however, due to the open-ended nature of human conversations, the quality of training data varies . |
| Approach: | They propose a data manipulation framework to augment and highlight effective training samples . they also propose to increase its manipulation skills through gradient descent with validation samples a reshaping framework to proactively restructure the data distribution towards reliable samples is also proposed . |
| Outcome: | The proposed framework improves the performance of open-domain neural dialogue models with respect to evaluation metrics and human judgments. |
Copied to clipboard
| Challenge: | Existing evaluation metrics for natural language generation are inadequate . existing metrics are not robust against simple perturbations and disagree with scores assigned by humans to perturbed output. |
| Approach: | They propose to propose checks which perturb the output and target a specific criteria and then use them to refine their evaluation. |
| Outcome: | The proposed templates show that existing evaluation metrics are not robust against simple perturbations and disagree with human scores on the perturbed output. |
Copied to clipboard
| Challenge: | Recent studies show that evaluating NLG systems using pairwise comparisons is expensive as the number of human annotations grows linearly with k. |
| Approach: | They propose a framework to efficiently identify the top-ranked system by actively choosing system pairs for comparison using dueling bandit algorithms. |
| Outcome: | The proposed framework reduces human annotations by 80% on 13 NLG evaluation datasets spanning 5 tasks . |
Copied to clipboard
| Challenge: | Recent studies have revealed that reading comprehension (RC) systems learn to exploit annotation artifacts and other biases in current datasets. |
| Approach: | They propose a task that requires giving answers and derivations to evaluate RC systems' internal reasoning. |
| Outcome: | The proposed framework annotates 4.6k questions with 3 reference derivations and shows that it is reliable and compares with existing benchmarks. |
Copied to clipboard
| Challenge: | Recent advances in conversational AI have been substantial, but developing real-time tasks guidance systems remains a challenge. |
| Approach: | They propose a data curation pipeline that synthesizes dialogues from annotated egocentric videos and a suite of automatic evaluation metrics that validated through extensive human studies. |
| Outcome: | The proposed framework synthesizes dialogues from annotated egocentric videos and validates them through extensive human studies. |
Copied to clipboard
| Challenge: | In order to evaluate large language models (LLMs), it is important to collect benchmark datasets in order to assess their multilingual performance. |
| Approach: | They extend the WMT24 dataset to cover 55 languages by collecting new human-written references and post-edits for 46 new languages/dialects. |
| Outcome: | The proposed dataset covers 55 languages and provides best-performing MT systems in all 55 languages. |
Copied to clipboard
| Challenge: | Generative large language models generate a high-dimensional probability distribution over all tokens in their vocabulary. |
| Approach: | They conduct extensive sensitivity analyses to determine how hyperparameter choices shape the outputs of generative large language models. |
| Outcome: | The proposed methods influence the distribution of diversity and coherence metrics in human-written text, but the optimal configurations vary across models and tasks. |
Copied to clipboard
| Challenge: | FarExStance is a new dataset for explainable stance detection in Farsi . it contains extractive explanations as evidence for stance labels and claims . |
| Approach: | They propose a dataset for explainable stance detection in Farsi with extractive explanations as evidence. |
| Outcome: | The proposed model is the most accurate on stance detection, while the best explanation is from few-shot Claude-3.5-Sonnet. |
Copied to clipboard
| Challenge: | Existing approaches to automate the complex task of translation are tedious and expensive. |
| Approach: | They describe acquisition, preprocessing, segmentation, and alignment of an Amharic-English parallel corpus. |
| Outcome: | The proposed corpus outperforms statistical machine translation models by six to seven BLEU points . the results show that the subword models outperformed word-based models by three to four BLUE points compared with the word-base models . |
Copied to clipboard
| Challenge: | Existing evaluation metrics for MCQ generation focus on the n-gram based similarity of the generated MCq to the gold sample and disregard their educational value. |
| Approach: | They propose to use a human survey to measure the MCQ’s answerability given knowledge of the target fact. |
| Outcome: | The proposed methods measure the MCQ’s answerability given knowledge of the target fact. |
Copied to clipboard
| Challenge: | Existing image captioning metrics are vulnerable to lexical perturbations, but they are not robust to such perturbations. |
| Approach: | They propose a perturbation-robust multilingual CLIPScore which is a reference-free image captioning metric for multiple languages. |
| Outcome: | The proposed metric outperforms baseline metrics in capturing lexical noise of all various perturbation types in all five languages while maintaining a strong correlation with human judgments. |
Copied to clipboard
| Challenge: | a recent data sampling method skews the annotated data toward shorter documents, not necessarily representative of the full test set. |
| Approach: | They examine different approaches to human evaluation and ranking of machine translation systems at the conference on machine translation . they propose a method that uses available labour budget to sample data in a more representative manner . |
| Outcome: | The proposed method improves representation of document lengths and produces stable rankings of translation quality. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have been adopted to process textual task description and accomplish procedural planning in embodied AI tasks because of their powerful reasoning ability. |
| Approach: | They propose to evaluate the planning ability of large language models and multi-modal counterfactual vision language models (VLMs) using a multi-factual household activity simulator and a chatGPT task description to evaluate their reasoning ability. |
| Outcome: | The proposed benchmark evaluates the planning ability of multi-modal and counterfactual vision language models on a household activity simulator and a chatGPT task description. |
Copied to clipboard
| Challenge: | Large language models are often defaulted to passive responses or narrow clarifications when faced with incomplete or under-specified prompts. |
| Approach: | They propose a new task paradigm where LLMs must identify gaps in context and strategically elicit implicit user knowledge through targeted questions. |
| Outcome: | The proposed framework outperforms o3-mini on evaluation metrics and human annotators favor clarification questions and final outlines. |
Copied to clipboard
| Challenge: | CR is the ability to understand and navigate the world using basic knowledge and understanding shared by most people. |
| Approach: | They propose to incorporate pretrained knowledge into NMT models and use them as robust testbeds for investigating CR in NMT. |
| Outcome: | The proposed method improves the training of NMT models with high CR abilities and provides accurate evaluation metrics. |
Copied to clipboard
| Challenge: | Existing approaches to evaluate open domain dialogues have a one-to-many problem . existing approaches lack commonsense reasoning biases and perform poorly in domain-specific scenarios. |
| Approach: | They propose a framework that leverages both a small, specialised model and LLMs for the evaluation of open-domain dialogues. |
| Outcome: | The proposed framework achieves state-of-the-art performance in both classification and evaluation tasks and exhibits better correlation with human judgements. |
Copied to clipboard
| Challenge: | a number of automated evaluation metrics are evaluated by multiple quality criteria, such as relevance, consistency, fluency and coherence. |
| Approach: | They propose a method that removes the confounding variable and detects unreliable correlations. |
| Outcome: | The proposed method detects unreliable correlations between QCs and human scores . it is based on a multi-QC setup, but it fails to detect summary corruptions . |
Copied to clipboard
| Challenge: | Interpolation-based retrieval-augmented language models (LMs) are a subtype of retrieval augmented language model that computes the probability of the next token by interpolating between the softmax distribution of the original LM and a token distribution formed by retrieving over an external datastore. |
| Approach: | They propose to interpolate the predicted distribution of the next word with a distribution formed from the most relevant retrievals for a given prefix. |
| Outcome: | The proposed methods do not exhibit improvements in open-ended generation quality, as measured by automatic evaluation metrics and human evaluations. |
Copied to clipboard
| Challenge: | a limited number of human annotations are required to evaluate multilingual summarization evaluation metrics. |
| Approach: | They propose a multilingual meta-evaluation framework that uses machine translation systems to transform a monolingual metaevaluations dataset into multilingual versions. |
| Outcome: | The proposed framework outperforms classical text-matching-based metrics in non-English languages. |
Copied to clipboard
| Challenge: | Existing metrics for text summarisation have restrictive token limits, limiting their effectiveness. |
| Approach: | They propose a human-annotated data set for evaluating automatic factuality metrics . they propose 'longDocFACTScore' framework which can be extended to any length document . |
| Outcome: | The proposed framework outperforms state-of-the-art metrics in evaluating long document summarisation data sets. |
Copied to clipboard
| Challenge: | Automatic Text Simplification (ATS) is a major natural language processing task that aims to help people understand complex text. |
| Approach: | They propose to use a human-annotated dataset to study automatic text simplification models to determine which metrics to use when evaluating new models. |
| Outcome: | The proposed models reconstruct the text into a simpler format by deletion, substitution, addition or splitting, while preserving the original meaning and correct grammar. |
Copied to clipboard
| Challenge: | Recent work has cast doubt on whether context-aware machine translation models learn useful signals from context or are improvements in automatic evaluation metrics just a side-effect. |
| Approach: | They propose to use separate encoders for source sentence and context as multiple sources for one target sentence to train context-aware machine translation models. |
| Outcome: | The proposed model improves translation quality even with empty lines as context, but the correct context improves it and random out-of-domain context degrades it. |
Copied to clipboard
| Challenge: | Large Language Models generate human-like text, making them unreliable for biomedical relation extraction tasks. |
| Approach: | They propose to use Large Language Models as judges to evaluate biomedical relation extraction . they propose structured output formatting for LLM-generated responses that helps LLMs improve their performance by 15%. |
| Outcome: | The proposed method improves LLM-Judges' performance by 15% . it is cheaper and more efficient than human evaluation metrics, the authors say . |
Copied to clipboard
| Challenge: | Current writing agents rely on predefined workflows and rigid thinking patterns to generate outlines before writing . authors propose a framework for long-form writing agents built on heterogeneous recursive planning . |
| Approach: | They propose a general agent framework that achieves human-like adaptive writing . they propose recursive task decomposition and dynamic integration of task types . |
| Outcome: | The proposed framework outperforms state-of-the-art approaches on both fiction and technical report generation. |
Copied to clipboard
| Challenge: | Existing automated evaluation metrics like ROUGE and BLEU show low correlation with human judgments. |
| Approach: | They propose a multi-agent evaluation framework that integrates multiple agents . they use ROUGE and BLEU to evaluate natural language models . |
| Outcome: | The proposed evaluation framework outperforms the current state-of-the-art methods in two meta-evaluation benchmarks. |
Copied to clipboard
| Challenge: | et al., 2010) show that hub embeddings are close to many unrelated examples in high-dimensional embeddable spaces . cross-modal encoders that project different modalities into a shared space are useful for cross-module applications . |
| Approach: | They propose a method for identifying the hub embedding and its corresponding hub text . they use images to evaluate cross-modal encoders that project different modalities into a shared space . |
| Outcome: | The proposed method can identify a single hub embedding and its corresponding hub text . it achieves comparable or higher similarity scores than human-written reference captions in many images . |